The healthcare sector generates an enormous volume of data, making its management and analysis a significant challenge. Machine learning techniques have emerged as effective tools for processing and interpreting such large-scale medical datasets. According to the National Family Health Survey (NFHS-5), the prevalence of thyroid disorders is steadily increasing in India, with approximately one in every ten adults affected. It is estimated that over 42 million people in the country are living with thyroid-related diseases. Accurate analysis of clinical data is essential for the timely and precise diagnosis of these conditions. In this study, thyroid disorders were categorized into five classes: Euthyroid, Overt Hypothyroidism, Overt Hyperthyroidism, Subclinical Hyperthyroidism, and Subclinical Hypothyroidism. The primary objective of this research is to perform an exploratory data analysis (EDA) of the thyroid dataset to uncover meaningful patterns and relationships among hormone levels, age, and gender. The findings from the EDA highlight significant gender-based differences, age-related variations, and important correlations among thyroid hormone parameters, providing valuable insights for thyroid disease assessment.
Introduction
Thyroid disorders are among the most common endocrine diseases in India and worldwide, affecting around 300 million people globally, including 42 million Indians. Women are 5–8 times more likely to develop thyroid disorders than men. Although thyroid diseases are widespread, fewer than 5% of cases are thyroid cancer.
The literature review highlights the importance of Exploratory Data Analysis (EDA) in thyroid disease diagnosis. Previous studies used datasets such as the UCI Thyroid Dataset to analyze clinical features, identify key biomarkers like TSH, handle missing data, reduce redundant features, and improve the performance of machine learning models. EDA helps uncover relationships among demographic, clinical, and laboratory variables, leading to highly accurate diagnostic systems.
The methodology used a dataset of 1,294 patient records with 8 attributes: Gender, Age, Location (Urban/Rural), T3, T4, TSH, Free T4, and Target (thyroid subtype). Data preprocessing was performed before analysis.
The Exploratory Data Analysis (EDA) included:
Univariate analysis to summarize individual variables.
Target distribution analysis, showing a nearly balanced distribution among thyroid disease subtypes, with Euthyroid being slightly more common.
Urban–rural analysis, indicating higher counts of certain thyroid subtypes in rural areas.
Age distribution, revealing that most thyroid cases occur between 30 and 50 years.
Density analysis, showing that TSH, T3, T4, and FT4 values are left-skewed.
Multivariate analysis to study relationships among multiple variables.
Age-wise analysis, showing different thyroid subtypes are most common between 25 and 55 years, with specific conditions concentrated in particular age ranges.
Gender-wise analysis, demonstrating that females are significantly more affected than males across all thyroid disease subtypes.
Conclusion
Since it is the most important thing to keep data set ready for the application of classification algorithm training available irrespective of the size of information and data mining skills, if the researcher does not make sense of data collection, a machine would be almost useless or even harmful. The exploratory analysis of thyroid dataset revealed compelling insights into the relationships between hormone levels, age, and gender. This phase sets the foundation for subsequent analyses and provides valuable insights into thyroid function, paving the way for further investigation and clinical implications. Key insights derived from the EDA is emphasizing gender-based disparities, age-related trends, and interrelationships among thyroid hormones.
References
[1] Abed, S., Saji, S., & Alshayeji, M. H., \"Predictive Analysis of Early Thyroid Disorders Using Integration of Data Mining and Ensemble Intelligence Approaches\", International Journal of Intelligent Systems,2026.
[2] Alam, M. Z., Rahman, R., Sozib, H. M., H. A., A. H., & Sabeena, A. A., “Enhancing Thyroid Disease Diagnosis With Machine Learning and Counterfactual Explainable AI”, IEEE, 2026.
[3] Hassan, A., Ramzan , S., Raza, A., Iqbal, M. M., Smerat , A., Fitriyani, N. L., . . . Lee, S. W. , “Improving thyroid disorder diagnosis via innovative stacking ensemble learning model”, Digit Health, 2025.
[4] Islam , S. S., Haque, M., Miah , M. U., Sarwar, T. B., & Nugraha , R., “Application of machine learning algorithms to predict the thyroid disease risk: an experimental comparative study”, PeerJ Comput Sci, 2022.
[5] Moharekar, D. T., Vadar, M. S., Pol, D. R., Bhaskar, D. C., & Moharekar , M. J., “Thyroid Disease Detection Using Machine Learning and Pycaret”, Specialusis Ugdymas / Special Education, 2022.
[6] Mollica , G., Francesconi , D., Costante , G., Moretti , S., Giannini, R., Puxeddu, E., & Valigi, P., “Classification of Thyroid Diseases Using Machine Learning and Bayesian Graph Algorithms”, IFAC-PapersOnLine. Elsevier, 2022.
[7] Osasere , O., Iqbal, M. Z., & Xining (Ning) , W., “Enhanced Diagnosis of Thyroid Diseases Through Advanced Machine Learning Methodologies”, Sci, 2025.
[8] S I TA L D I N, N. “Thyroid Disease Prediction and Symptom Exploration During Early Pregnancy”, 2022
[9] Thalor, M., Pathak, M., Kale, V., & Bhende, V., “Automated Diagnosis to Predict the Thyroid Using Machine Learning Algorithms”, SSRG International Journal of Electrical and Electronics Engineering, 2024.
[10] Tiwari, Y., Saxena, D., Pal, P., Dixit, H. M., & Yadav, A., “Thyroid Disease Prediction: Leveraging Machine Learning for Early Recurrence Detection”, 12th International Conference on Computing for Sustainable Global Development (INDIACom), Delhi, India: IEEE, 2025